Papers with quality assurance
Synthetic Data for Evaluation: Supporting LLM-as-a-Judge Workflows with EvalAssist (2025.emnlp-demos)
Copied to clipboard
Martín Santillán Cooper, Zahra Ashktorab, Hyo Jin Do, Erik Miehling, Werner Geyer, Jasmina Gajcin, Elizabeth M. Daly, Qian Pan, Michael Desmond
| Challenge: | EvalAssist is a web-based application designed to assist human-centered evaluation of language model outputs. |
| Approach: | They propose a synthetic data generation tool integrated into EvalAssist to assist human-centered evaluation of language model outputs. |
| Outcome: | The proposed tool supports flexible prompting, RAG-based grounding, persona diversity, and iterative generation workflows. |
PromptLab: A Collaborative Platform for Prompt Engineering and Dataset Curation (2026.eacl-demo)
Copied to clipboard
Maged S. Al-shaibani, Zaid Alyafeai, Dania Refai, Nawaf Alomari, Ahmed Ashraf, Mais Alheraki, Mustafa Alturki, Hamzah Luqman, Irfan Ahmad
| Challenge: | PromptLab is a web-based prompt engineering platform for collaborative prompt development across diverse natural language processing tasks and datasets. |
| Approach: | They propose to integrate prompt generation via OpenRouter and provide real-time validation with multiple Large Language Models. |
| Outcome: | The platform addresses primary challenges in prompt development, including template creation, collaborative review, and quality assurance through a comprehensive workflow that supports both individual researchers and team-based projects. |
Sensing and Learning Human Annotators Engaged in Narrative Sensemaking (N18-4)
Copied to clipboard
| Challenge: | a substantial sector of the gig economy is the use of crowdworkers to annotate data for machine learning and analysis. |
| Approach: | They propose a narrative-sorting annotation task that sorts tweets chronologically by topic, emotional content, and length. |
| Outcome: | The proposed task enables readers to sort sequential, target-topical, and emotionally emotional tweets. |
AI Coach Assist: An Automated Approach for Call Recommendation in Contact Centers for Agent Coaching (2023.acl-industry)
Copied to clipboard
| Challenge: | In recent years, the utilization of Artificial Intelligence (AI) in the contact center industry is on the rise. |
| Approach: | They present a transformer-based pairwise sentence classification model that analyzes call transcripts to determine which calls are most relevant for coaching purposes. |
| Outcome: | The proposed model can determine which calls are most relevant for coaching purposes based on quality assurance queries/questions asked by managers or supervisors . |
Lost and Found: Computational Quality Assurance of Crowdsourced Knowledge on Morphological Defectivity in Wiktionary (2025.acl-srw)
Copied to clipboard
| Challenge: | a recent study shows that wikis are not reliable for linguistic knowledge of defects in understudied languages. |
| Approach: | They customize a neural morphological analyzer to annotate Latin and Italian corpora . they validated morphology using crowd-sourced data from Wiktionary to find defects . |
| Outcome: | The proposed algorithm annotates Latin and Italian corpora using crowd-sourced data . results show that 7% of Latin lemmata listed as defective show strong corpus evidence of being non-defective. |
Translation Crowdsourcing: Creating a Multilingual Corpus of Online Educational Content (L18-1)
Copied to clipboard
Vilelmini Sosoni, Katia Lida Kermanidis, Maria Stasimioti, Thanasis Naskos, Eirini Takoulidou, Menno van Zaanen, Sheila Castilho, Panayota Georgakopoulou, Valia Kordoni, Markus Egg
| Challenge: | a large corpus of online content has been developed via large-scale crowdsourcing. |
| Approach: | They describe a multilingual corpus of online content that has been manually translated into 11 European and BRIC languages using the crowdsourcing platform. |
| Outcome: | The proposed corpus is a product of the EU-funded TraMOOC project and is used to train, tune and test machine translation engines. |
Experience Report: Implementing Machine Translation in a Regulated Industry (2025.emnlp-industry)
Copied to clipboard
| Challenge: | a global medical technology company has invested substantial resources in translating content into the various languages required across their global markets. |
| Approach: | They propose to use human-in-the-loop validation to evaluate machine translation systems in a medical technology company. |
| Outcome: | The proposed method dominates reviewer preference across all languages and tones of interest, the authors show . the "Gold" control ranks poorly in one language and the lower ranks have high variance. |
Self-prompted Chain-of-Thought on Large Language Models for Open-domain Multi-hop Reasoning (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing open-domain question-answering methods lack quality assurance . existing methods lack scalability and poor diversity, hindering LLMs' capabilities . |
| Approach: | They propose an open-domain multi-hop reasoning framework to answer multi-choice questions . they propose an adaptive sampler for in-context selection and self-prompted inference . |
| Outcome: | The proposed framework surpasses the existing SOTA methods on large-scale datasets and doubles the zero-shot performance of small-scale LLMs. |
BertNet: Harvesting Knowledge Graphs with Arbitrary Relations from Pretrained Language Models (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing methods to construct knowledge graphs are limited to a small set of relations due to manual cost or restrictions in text corpus. |
| Approach: | They propose to automatically construct knowledge graphs (KGs) of diverse new relations from pretrained language models that accept knowledge queries with prompts. |
| Outcome: | The proposed framework extracts knowledge of over 400 new relations from pretrained language models, including RoBERTaNet, with minimal input of a relation definition and a few shot of example entity pairs. |
FRAME: Feedback-Refined Agent Methodology for Enhancing Medical Research Insights (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing approaches to automate scientific research are limited by human cognitive constraints and timeintensive workflows. |
| Approach: | They propose a framework that enhances medical paper generation through iterative refinement and structured feedback. |
| Outcome: | The proposed framework achieves significant improvements over conventional methods across multiple models and evaluation dimensions. |
Intrinsic Evaluation of Summarization Datasets (2020.emnlp-main)
Copied to clipboard
| Challenge: | Almost all popular summarization datasets do not come with inherent quality assurance guarantees. |
| Approach: | They propose to use 5 metrics to evaluate quality of summarization datasets . they find that data usage in recent summarizing research is inconsistent with the properties of the data. |
| Outcome: | The proposed metrics can be inexpensive heuristics for detecting generically low quality examples. |
PMIndiaSum: Multilingual and Cross-lingual Headline Summarization for Languages in India (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing datasets for Indian languages are limited in terms of coverage and size. |
| Approach: | They propose a multilingual and massively parallel summarization corpus focused on languages in India that provides a training and testing ground for four language families, 14 languages, and the largest to date with 196 language pairs. |
| Outcome: | The proposed dataset provides a training and testing ground for four language families, 14 languages, and the largest to date with 196 language pairs. |
LongMP-Bench: A Benchmark for Multimodal Persona Understanding in Long-Term Dialogues (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing datasets suffer from limited persona diversity and static, overly simplified settings, making them insufficient for capturing the complexity of real-world interactions. |
| Approach: | They propose a benchmark to evaluate models' ability to understand evolving user personas within long-term multimodal dialogues by using a dataset that contains long conversations from 150 users. |
| Outcome: | The proposed benchmark aims to assess models' ability to track persona evolution, integrate visual and textual inputs, and apply persona understanding in realistic dialogue scenarios. |